Papers with annotation guidelines

55 papers
PharmaCoNER: Pharmacological Substances, Compounds and proteins Named Entity Recognition track (D19-57)

Copied to clipboard

Challenge: Biomedical text mining is one of the most prolific application domains of natural language processing technologies.
Approach: They propose to share a task on detecting drug and chemical entities in medical documents in Spanish with other languages to improve access to biomedical text mining.
Outcome: The first task on detecting drug and chemical entities in Spanish medical documents yielded competitive results with F-measures above 0.91.
LENS: Learning Entities from Narratives of Skin Cancer (2025.coling-demos)

Copied to clipboard

Challenge: Learning entities from narratives of skin cancer (LENS) is an automatic entity recognition system built on colloquial writings from skin cancer-related forums.
Approach: They propose to use reddit forums to create an automatic entity recognition system that can be used to predict skin cancer outcomes.
Outcome: LENS achieves an overall entity-level F1 score of 0.561 . other notable results include “CANC_T” (0.747), “STG” (0.888), “POB” (0.914), “GENDER” (0.750), “A/G” (00.646), “EMO” (0.619), and “MHD” (0.503).
Automatic Focus Annotation: Bringing Formal Pragmatics Alive in Analyzing the Information Structure of Authentic Data (N18-1)

Copied to clipboard

Challenge: Using focus-background dichotomy, discourse and information structure of sentences are being studied in context.
Approach: They propose to automate the analysis of focus in authentic written data by using a range of lexical, syntactic, and semantic features to achieve an accuracy of 78.1%.
Outcome: The proposed approach achieves 78.1% accuracy for identifying focus in authentic written data.
GATE Teamware 2: An open-source tool for collaborative document classification annotation (2023.eacl-demo)

Copied to clipboard

Challenge: GATE Teamware 2 is an open-source web-based platform for managing teams of annotators working on document classification tasks.
Approach: They present GATE Teamware 2: an open-source web-based platform for managing teams of annotators working on document classification tasks.
Outcome: GATE Teamware 2 is an open-source web-based platform for managing teams of annotators working on document classification tasks.
A Corpus for Sentence-Level Subjectivity Detection on English News Articles (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to spotting subjectivity require language-specific tools.
Approach: They develop annotation guidelines for sentence-level subjectivity detection that are not limited to language-specific cues.
Outcome: The proposed framework enables subjectivity detection in English and across other languages without relying on language-specific tools, such as lexicons or machine translation.
Chop and Change: Anaphora Resolution in Instructional Cooking Videos (2022.findings-aacl)

Copied to clipboard

Challenge: temporally evolving entities present challenges for anaphora resolution tasks . recipes provide rich source for referring expressions of transformed entities .
Approach: They propose to use annotations to annotate recipes for anaphora resolution task . they propose to employ temporal features to improve anamorphic resolution .
Outcome: The proposed annotation scheme improves the performance of the anaphora resolution task.
ViHOS: Hate Speech Spans Detection for Vietnamese (2023.eacl-main)

Copied to clipboard

Challenge: Increasing use of social networking sites can cause problems for human moderators to review tagged comments.
Approach: They present a dataset that contains 26k spans on 11k comments and detailed annotation guidelines . they also provide definitions of hateful and offensive spans in Vietnamese comments .
Outcome: The proposed dataset shows that it is difficult to detect specific types of spans in the dataset . the dataset is the first human-annotated corpus containing 26k spans on 11k comments .
MIMICause: Representation and automatic extraction of causal relation types from clinical notes (2022.findings-acl)

Copied to clipboard

Challenge: Extracted causal information from clinical notes can be combined with structured EHR data such as demographics, diagnoses, and medications.
Approach: They propose to annotate clinical notes and develop an annotated corpus and provide baseline scores to identify types and direction of causal relations between a pair of biomedical concepts.
Outcome: The proposed annotation guidelines achieved a high inter-annotator agreement and a macro F1 score on the clinical text.
RuSentiment: An Enriched Sentiment Analysis Dataset for Social Media in Russian (C18-1)

Copied to clipboard

Challenge: RuSentiment is currently the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post).
Approach: They propose to use RuSentiment to annotate social media posts in Russian with a kappa of 0.58 and a set of annotation guidelines that are extensible to other languages.
Outcome: The proposed dataset is the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post).
A French Medical Conversations Corpus Annotated for a Virtual Patient Dialogue System (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for creating virtual patient dialogue systems require large data specific to the language, domain and clinical cases studied.
Approach: They propose to build an annotated corpus of medical dialogues in french using medical interviews and a data annotation scheme.
Outcome: The proposed corpus is made publicly available under a Free/Libre Open Source licence.
A Corpus for Argumentative Writing Support in German (2020.coling-main)

Copied to clipboard

Challenge: In today's world most information is readily available. Consequently, the sole reproduction of information is losing attention.
Approach: They propose an annotation approach to capture claims and premises of arguments and their relations in student-written peer reviews on business models in german language.
Outcome: The proposed annotation scheme guides annotators to moderate agreement with the proposed scheme on 50 persuasive student-written peer reviews on business models.
Annotating Spin in Biomedical Scientific Publications : the case of Random Controlled Trials (RCTs) (L18-1)

Copied to clipboard

Challenge: Fig. 6: Annotation of biomedical abstracts for automatic detection of inadequate claims (spin) spin is a misleading presentation of scientific results in randomized controlled trials, an important type of clinical trial.
Approach: They propose an algorithm for automatic detection of inadequate claims (spin) they propose to use a corpus of biomedical articles for the task .
Outcome: The proposed algorithm can detect inadequate claims in biomedical abstracts without requiring any prior knowledge of the literature.
Towards a Gold Standard Corpus for Variable Detection and Linking in Social Science Publications (L18-1)

Copied to clipboard

Challenge: a new corpus for detecting and linking survey variables is being developed . the corpus is multilingual and includes manually curated word and phrase alignments .
Approach: They propose to create a corpus for the evaluation of detecting and linking survey variables in social science publications.
Outcome: The proposed corpus is the first gold standard for the variable detection and linking task.
ParCorFull2.0: a Parallel Corpus Annotated with Full Coreference (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpus ParCorFull contains parallel texts for English-German, French and Portuguese . translation of coreference across languages is challenging for MT and other NLP applications .
Approach: They describe a parallel corpus annotated with full coreference chains for multiple languages . they use the existing corpus ParCorFull to study translation of coreference across languages - a challenge for machine translation and NLP .
Outcome: The proposed corpus addresses translation of coreference across languages, a problem still challenging for machine translation and other multilingual natural language processing applications.
Automatic Section Recognition in Obituaries (2020.lrec-1)

Copied to clipboard

Challenge: Obituaries contain information about people’s values across times and cultures, which makes them useful for exploring cultural history.
Approach: They propose to use a convolutional neural network to recognize these sections in obituaries to improve their annotation.
Outcome: The proposed model outperforms bag-of-words and embedding-based BiLSTMs and BiLStm-CRFs with a micro F1 = 0.81.
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)

Copied to clipboard

Challenge: In literature, spoken interactions between characters are of central importance to the narrative.
Approach: They propose to annotate quotations, including their interpersonal structure, for English literary text.
Outcome: The proposed dataset provides a rich view of dialogue structures not available from other available corpora.
Arap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification (L18-1)

Copied to clipboard

Challenge: Existing corpus of Arabic textual data is limited to English or other European languages.
Approach: They present a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the arab world representing the major Arabic dialectal varieties.
Outcome: The provided corpus will enrich the limited set of available language resources for Arabic and be invaluable enabler for developing author profiling tools and NLP tools for Arabic.
The SOFC-Exp Corpus and Neural Approaches to Information Extraction in the Materials Science Domain (2020.acl-main)

Copied to clipboard

Challenge: Using BERT embeddings leads to large performance gains, but with increasing task complexity, adding a recurrent neural network seems beneficial.
Approach: They propose an annotation scheme for marking information on publications related to solid oxide fuel cells . they propose to use a recurrent neural network to solve a variety of tasks .
Outcome: The proposed scheme is based on a corpus of 45 open-access scholarly articles and a neural network for a variety of tasks.
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs (2023.findings-eacl)

Copied to clipboard

Challenge: Existing graph-alignment metrics that measure graph distances are not reliable, we show . metric is spread out and does not provide upper bounds for extended tasks.
Approach: They propose a metric to measure a distance between graphs by aligning nodes and counting matching graph triples.
Outcome: The proposed method reduces search space and improves scoring by reducing the number of errors.
MuCPAD: A Multi-Domain Chinese Predicate-Argument Dataset (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies show that shallow semantic role labeling (SRL) performance drops under out-of-domain setting.
Approach: They propose to annotate a multi-domain Chinese predicate-argument dataset using a frame-free annotation methodology and strict double annotation for improving data quality.
Outcome: The proposed dataset is compared with a dataset from six different domains.
Structured Persuasive Writing Support in Legal Education: A Model and Tool for German Legal Case Solutions (2023.findings-acl)

Copied to clipboard

Challenge: An annotation approach for capturing structured components and arguments in legal case solutions of German students is proposed based on the appraisal style, which dictates the structured way of persuasive writing in law.
Approach: They propose an annotation scheme with annotation guidelines that identify structured writing in legal case solutions.
Outcome: The proposed approach captures the structure of a persuasive legal text and can be predicted using transformer-based models.
From Multiple-Choice to Extractive QA: A Case Study for English and Arabic (2025.coling-main)

Copied to clipboard

Challenge: Recent years have brought about very fast developments in Natural Language Processing (NLP), but many other languages are overlooked due to limited resources.
Approach: They propose to repurpose a multilingual BELEBELE dataset for a task of extractive QA in the style of machine reading comprehension.
Outcome: The proposed approach could be used to extract QA in the style of machine reading comprehension.
Annotating Arguments in a Corpus of Opinion Articles (2022.lrec-1)

Copied to clipboard

Challenge: Argument annotation is the process of exposing and justifying one's points of view, with the aim of conveying a logical reasoning through a set of semantically related propositions.
Approach: They propose to use argumentative discourse units to annotate arguments in Portuguese using a multi-layered process to analyze the annotations produced.
Outcome: The proposed model exploits the best practices identified in previous studies while fostering the potential use of the resulting annotated corpus for new purposes.
MuCGEC: a Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction (2022.naacl-main)

Copied to clipboard

Challenge: Using a multi-reference multi-source evaluation dataset, Chinese grammatical error correction (CGEC) is relatively scarce.
Approach: They propose a multi-reference multi-source evaluation dataset for Chinese grammar error correction . the dataset contains 7,063 sentences written by Chinese-as-a-Second-Language learners .
Outcome: The proposed dataset can be used to evaluate Chinese grammar errors in Chinese.
The Causal News Corpus: Annotating Causal Relations in Event Sentences from News (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotation guidelines for event causality focus on only explicit relations or clauses.
Approach: They propose an annotation schema for event causality that addresses these concerns . they annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not.
Outcome: The proposed annotation schema for event causality addresses these concerns . it performs well with 81.20% F1 score on test set and 83.46% in 5-folds cross-validation .
What Did You Learn To Hate? A Topic-Oriented Analysis of Generalization in Hate Speech Detection (2023.eacl-main)

Copied to clipboard

Challenge: Hate speech detection datasets often use different annotation guidelines, resulting in inconsistencies . authors propose a topic-oriented approach to study generalization across popular hate speech datasets .
Approach: They propose a topic-oriented approach to study generalization across popular hate speech datasets . they compare Transformer-based models in capturing topic-generic and topic-specific knowledge .
Outcome: The proposed approach improves the reliability of hate speech detection on social media platforms.
LPAttack: A Feasible Annotation Scheme for Capturing Logic Pattern of Attacks in Arguments (2022.lrec-1)

Copied to clipboard

Challenge: Argumentation plays a central role in human communication, where refuting or attacking others’ arguments is a common persuasion strategy.
Approach: They propose a novel annotation scheme that captures common modes and complex rhetorical moves in attacks along with the implicit presuppositions and value judgments.
Outcome: The proposed scheme shows moderate agreement between the two annotations, indicating that human annotation is feasible.
Developing the Bangla RST Discourse Treebank (L18-1)

Copied to clipboard

Challenge: a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla .
Approach: They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus .
Outcome: The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications .
Supporting Cognitive and Emotional Empathic Writing of Students (2021.acl-long)

Copied to clipboard

Challenge: Empathy skills are an elementary skill in society for daily interaction and professional communication and are therefore elementary for educational curricula.
Approach: They propose an annotation approach to capture emotional and cognitive empathy in student-written peer reviews on business models in germany.
Outcome: The proposed annotation scheme guides annotators to a substantial to moderate agreement with the model and shows that it is effective.
Hype or not? Formalizing Automatic Promotional Language Detection in Biomedical Research (2026.eacl-long)

Copied to clipboard

Challenge: Promotional language is a term used to undermine objective evaluation of evidence, impede research development, and erode trust in science.
Approach: They propose formalized guidelines for identifying hype language and apply them to annotate a portion of the National Institutes of Health grant application corpus.
Outcome: The proposed guidelines can help humans reliably annotate candidate hype adjectives and train machine learning models yield promising results.
Constructing a Dependency Treebank for Second Language Learners of Korean (2024.lrec-main)

Copied to clipboard

Challenge: a manually annotated syntactic treebank is available for second language learners . the dataset includes 7,530 sentences (66,982 words; 129,333 morphemes)
Approach: They propose to manually annotate syntactic treebanks based on Universal Dependencies from Korean written data.
Outcome: The proposed dataset includes 7,530 sentences and 129,333 morphemes from Korean learners.
BKTreebank: Building a Vietnamese Dependency Treebank (L18-1)

Copied to clipboard

Challenge: In this paper, we present the building of a dependency treebank for Vietnamese .
Approach: They propose to build a Vietnamese dependency treebank using automatic taggers and automatic tagging.
Outcome: The proposed treebank is a useful resource for Vietnamese language processing.
Overlaps and Gender Analysis in the Context of Broadcast Media (2022.lrec-1)

Copied to clipboard

Challenge: Using gender and overlap annotations, we characterise interactions between speakers according to their gender and role in broadcast media.
Approach: They propose to characterise interactions between speakers according to their gender and role in broadcast media by using a small dataset of 93 recordings from LCP French channel.
Outcome: The proposed method could improve the efficiency of qualitative studies conducted in human sciences.
Translation and Fusion Improves Cross-lingual Information Extraction (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown significant progress in information extraction tasks due to lack of labeled data for fine-tuning and unlabeled text for pre-training.
Approach: They propose a framework in which large language models are fine-tuned to use English translations of low-resource language data.
Outcome: The proposed model improves cross-lingual transfer over the base model on 12 multilingual IE datasets spanning 50 languages.
RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced Labour (2022.lrec-1)

Copied to clipboard

Challenge: Forced labour is the most common type of modern slavery, affecting at least 24.9 million people worldwide.
Approach: They propose to annotate an English corpus for multi-class and multi-label forced labour detection using specialised data from specialised sources.
Outcome: The proposed corpus consists of 989 news articles annotated according to risk indicators defined by the International Labour Organization (ILO).
Morphological Segmentation for Low Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages .
Approach: This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program.
Outcome: The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications.
Building an English-Chinese Parallel Corpus Annotated with Sub-sentential Translation Techniques (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that human translators often resort to different non-literal translation techniques besides literal translation . however, they receive less attention in developing natural language processing (NLP) applications.
Approach: They propose to have a better semantic control of extracting paraphrases from bilingual parallel corpora.
Outcome: The proposed method can automatically recognize different non-literal translation techniques . the results confirm the hypothesis of the proposed method .
Automating Idea Unit Segmentation and Alignment for Assessing Reading Comprehension via Summary Protocol Analysis (2022.lrec-1)

Copied to clipboard

Challenge: In second language learning, summaries are among the most popular type of student assignments.
Approach: They propose to revise the annotation guidelines to allow machine implementation of the new annotation guidelines.
Outcome: The proposed algorithm achieves 0.789 precision and 0.844 recall over the L2WS 2021 corpus.
Modeling Persuasive Discourse to Adaptively Support Students’ Argumentative Writing (2022.acl-long)

Copied to clipboard

Challenge: Argumentation is an omnipresent rudiment of daily communication and thinking . humans struggle to develop argumentation skills due to a lack of individual and instant feedback in their learning process.
Approach: They propose an argumentation annotation approach to model argumentative discourse in student-written business model pitches and embed it into an adaptive writing support system for students that provides individual argumentation feedback.
Outcome: The proposed method annotates a corpus of 200 business model pitches in german and measures their self-efficacy and ease-of-use in a real-world writing exercise.
ODIL_Syntax: a Free Spontaneous Spoken French Treebank Annotated with Constituent Trees (2020.lrec-1)

Copied to clipboard

Challenge: ODIL Syntax is a French treebank built on spontaneous speech transcripts . the structure of every speech turn is represented by constituent trees .
Approach: They propose a French treebank built on spontaneous speech transcripts with a constituency tree representation.
Outcome: The proposed treebank is based on the French TreeBank, with some annotation guidelines . the proposed tree bank will be freely distributed by January 2020 under a Creative Commons licence .
From Laughter to Inequality: Annotated Dataset for Misogyny Detection in Tamil and Malayalam Memes (2024.lrec-main)

Copied to clipboard

Challenge: a new form of memes has emerged to combat misogyny and harmful stereotypes . authors present a dataset to analyze online misogamy in Tamil and Malayalam communities .
Approach: They propose to create an annotated dataset with detailed annotation guidelines to analyze online misogyny within Tamil and Malayalam-speaking communities.
Outcome: The proposed dataset reveals the world of gender bias and stereotypes in Tamil and Malayalam-speaking communities.
GERMS-AT: A Sexism/Misogyny Dataset of Forum Comments from an Austrian Online Newspaper (2024.lrec-main)

Copied to clipboard

Challenge: sexism/misogyny dataset extracted from comments of online forum of newspaper . corpus of 8 000 comments annotated with 5 levels of sexist/mistoginist .
Approach: They present a sexism/misogyny dataset extracted from comments of an online forum of an Austrian newspaper.
Outcome: The results show that the corpus of comments is sexist/misogynistic and has 5 levels of sexism/mistoginess.
Decomposing and Comparing Meaning Relations: Paraphrasing, Textual Entailment, Contradiction, and Specificity (2020.lrec-1)

Copied to clipboard

Challenge: SHARel is a new typology for decomposing and comparing multiple meaning relations . it consists of 26 linguistic and 8 reason-based categories and can be applied to all relations with a high inter-annotator agreement.
Approach: They propose a new typology that consists of 26 linguistic and 8 reason-based categories and propose SHARel for decomposing and comparing multiple meaning relations.
Outcome: The proposed method can be applied to all relations with high inter-annotator agreement.
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions.
Approach: They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language .
Outcome: The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language .
I Feel Offended, Don’t Be Abusive! Implicit/Explicit Messages in Offensive and Abusive Language (2020.lrec-1)

Copied to clipboard

Challenge: Recent literature suggests different approaches to identify abusive language phenomena . however, there is a lack of data sets that take into account the degree of explicitness .
Approach: They propose to use annotation guidelines to distinguish between explicit and implicit abuse in English and apply them to OLID/OffensEval.
Outcome: The proposed tool distinguishes between explicit and implicit abuse in English and takes into account the degree of explicitness.
A Manually Annotated Resource for the Investigation of Nasal Grunts (2020.lrec-1)

Copied to clipboard

Challenge: acoustic annotation of nasal grunts is described in the whole CID corpus of the french language . acculturation of non-lexical conversational sounds has been debated for a long time .
Approach: They propose an annotation framework for nasal grunts of the whole French CID corpus . they characterise acoustic cues and visual cue conventions followed for the annotation .
Outcome: The proposed framework is based on the entire French CID corpus.
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments (2023.acl-long)

Copied to clipboard

Challenge: Existing evaluations of human-in-the-loop systems to combat misinformation are often set up automatically using datasets that were retrospectively constructed.
Approach: They propose a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them.
Outcome: The proposed framework is based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments.
Analyzing Gambling Addictions: A Spanish Corpus for Understanding Pathological Behavior (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the interaction between natural language use and gambling disorders.
Approach: They build a new corpus of sentences that are searched and compared using top-k pooling to form the assessment pools of sentences.
Outcome: The proposed model is based on a new corpus of sentences in spanish .
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)

Copied to clipboard

Challenge: e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text.
Approach: They propose a timeline-based framework that achieves full coverage of all possible TLINKs.
Outcome: The proposed framework achieves full coverage of all possible TLINKs in a text.
TeClass: A Human-Annotated Relevance-based Headline Classification and Generation Dataset for Telugu (2024.lrec-main)

Copied to clipboard

Challenge: Relevance-based headline classification is under-explored in low-resource languages like Telugu due to a lack of annotated data.
Approach: They propose that relevance-based headline classification can greatly aid the task of generating relevant headlines.
Outcome: The proposed model can generate relevant headlines with 78,534 annotations in Telugu . the model shows a 5 point increment in the ROUGE-L scores .
AdabNER: Arabic Digital Archive Books with Nested Entity Recognition (2026.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of information extraction that classifies entities into predefined categories like person names.
Approach: They propose a large-scale nested Arabic Named Entity Recognition dataset . they fine-tuned five pre-trained Arabic BERT encoders in two settings .
Outcome: The first large-scale nested NER dataset for Arabic literary texts is published online . the dataset yields 78,530 entity mentions, 18.96% of which are nestated .
R1-RE: Cross-Domain Relation Extraction with RLVR (2026.acl-long)

Copied to clipboard

Challenge: Relation extraction (RE) is a core task in natural language processing.
Approach: They propose a supervised learning task for relation extraction (RE) based on annotation guidelines.
Outcome: The proposed model achieves an average OOD accuracy of 70%, on par with leading proprietary models such as GPT-4o.
Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale Recognition (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on a single type of quantity or a specific format, lacking a comprehensive evaluation of scale recognition capabilities.
Approach: They propose a visual scale recognition benchmark built using images from COCO, Open Images, and Flickr to evaluate scale recognition capabilities of multimodal large language models.
Outcome: The proposed model achieves 42.60% accuracy, lower than the 97.40% of humans.
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation (2026.findings-acl)

Copied to clipboard

Challenge: Existing MT evaluation frameworks fail to capture dialect- and culture-specific errors in diglossic languages.
Approach: They propose a hierarchical error taxonomy for diagnosing MT errors through six linguistic levels: sociolinguistics, pragmatics, semantics, morphosyntax, orthography, and graphetics.
Outcome: The proposed framework produces 6,113 labeled error spans across 3,495 unique erroneous sentences . it is language-agnostic and can be easily applied to or adapted for other languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations